Papers with Text-to-Speech models

2 papers
PresentAgent: Multimodal Agent for Presentation Video Generation (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing methods for generating static slides or text summaries are limited to producing narrated presentations.
Approach: They propose a multimodal agent that transforms long-form documents into narrated presentations.
Outcome: The present agent produces fully synchronized visual and spoken content that closely mimics human-style presentations.
LibriS2S: A German-English Speech-to-Speech Translation Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in speech-to-text translation have led to significant improvements, but the availability of appropriate training data is limiting.
Approach: They propose a new text-to-speech and speech-tospech translation model that directly learns to generate the speech signal based on the pronunciation of the source language.
Outcome: The proposed model learns to generate speech signal based on pronunciation of source language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations